This is the owning Tier-1 document for the audio domain. Until this version it was a two-line title with no body, and the domain's ruled canon lived scattered across proposal-tier program documents, technical dossiers, the decisions register and the pipeline ledger. The ledger names that absence directly: the C9 VO row records "T1_Audio_Spec [DRAFT v0.1] is a 2-line empty stub" as the canon-home gap (docs/PIPELINE_LEDGER.md §4 row C9), the F1 music row repeats it (§7 row F1), the F2 SFX row records "no T1 owner (T1_Audio_Spec stub)" (§7 row F2), and §11 carries lifting the stub as a named work order (§11, VETOABLE 3).
This version is CONSOLIDATION, not invention. Every bar, rule and posture below is lifted from an existing ruling, doctrine, dossier or registry and carries its citation. Where the sources contradict one another, the resolution and its reasoning are stated in the body and the residue is carried to §10 rather than smoothed over. Nothing genuinely new is absorbed silently into the body.
The document is authored to DRAFT. A fresh-context critic gates any DRAFT-to-ACTIVE flip.
§1 Authority and scope
1.1 What this document owns
This document owns the audio domain: what the game's music, sound effects, foley and voice must BE, the quality bars they are held to, the care discipline that governs them, and the evidence a claim about them has to carry. It owns the four audio surfaces as one domain, because the region-page template already treats them as one consumption unit and the doctrine, the rubric and the dossiers all cross-reference each other constantly (docs/translation/T99_Translation_Audio.md §1).
Music, including themes, leitmotifs, regional cues, beds, ambience and stingers.
Sound effects, including environmental ambient, magic-effect, creature vocalisation and foley.
Foley and footstep, as the wired sub-case of sound effects.
Voice and VO, including the natural-voice register every player-facing line carries.
1.2 Where it sits in the authority hierarchy
The CVD wins every conflict. Where this document and the CVD disagree, this document is the defect. The care standard of CVD §17.1 and the authoring-output standards of CVD §20 bind everything here.
T1_Build_Pipeline_Contracts [ACTIVE v1.0] owns the generation CONTRACT surface for the three audio pipelines: §3 the music pipeline contract, §4 the voice pipeline contract, §10 the SFX pipeline contract. That document owns foreign-key mapping, required columns and state-machine behaviour. This document owns the domain content those contracts carry. The two are complementary and neither overrides the other on the other's ground.
The runtime realization layer beneath both is docs/translation/T99_Translation_Audio.md, which says so on its own face: "It does not carry canonical authority of its own" (§1). It supplies mechanism — install lanes, import automation, MetaSounds and Quartz wiring — and never content canon.
The proposal-tier program documents named in §1.4 are HOW references under this document. They are subordinate to it, and to canon generally, by their own declared header discipline.
1.3 The T0 registries this document owns row-level canon for
T0_SFX_Registry [ACTIVE v1.0] — 680 rows across 41 columns, the second-largest registry in the corpus (docs/PIPELINE_LEDGER.md §7 row F2). Categories as populated: MAGIC_EFFECT 652, ENVIRONMENTAL_AMBIENT 20, FOLEY 7, CREATURE_VOCALIZATION 1. Classes as populated: gameplay_cue 658, ambient 17, cultural_traditional 4, cinematic_cue 1.
T0_Voice_Registry [DRAFT v0.1] — 23 rows across 21 columns: 13 cultural_register rows and 10 tts_casting rows. Every tts_casting row carries voice_clone_status synthetic_only.
T0_Motif_Registry [DRAFT v0.1] — 1 row across 28 columns. This is the sharpest data gap in the domain and the ledger names it as such: one motif row under a melody-first thirty-year bar (docs/PIPELINE_LEDGER.md §7 row F1).
1.4 The proposal-tier program documents this document governs
Each of these remains the working HOW for its lane and each is subordinate to this document.
The composition lane's runbook is docs/proposals/MUSIC_COMPOSITION_DOCTRINE.md: leitmotif system, hummability checklist, wordless-vocal register, fourteen per-culture register briefs for the slice, orchestration and the concert deliverable, adaptive architecture, the workflow, and the consolidated registry schema at Appendix S.
The per-chapter score program and the measurement science behind it is docs/proposals/music/MASTERPIECE_PROGRAM.md.
The roster, the families, the deny lists and the reveal firewall are docs/proposals/music/LEITMOTIF_ARCHITECTURE.md, applied to T0_Theme_Registry on 2026-08-05.
The exemplar-anchored scoring instrument is docs/proposals/music/NOSTALGIA_RUBRIC.md.
The per-theme constraint set and the twelve slice register cards are docs/proposals/music/COMPOSER_BRIEF_PACK.md.
The music and SFX lane's law for stack, licence and records is docs/pipeline_review/tech_research/PIPE_AUDIO_MUSIC_2026-07-29.md.
The VO lane's stack, licence and care-review contract is docs/pipeline_review/tech_research/PIPE_VOICE_2026-07-29.md.
The natural-voice register is docs/NATURAL_VOICE_DOCTRINE.md, binding on every player-facing word and therefore on every spoken line and every sung, chanted or diegetic text surface.
§2 The standing direction — the ruled bars
These six are Josh's own direction for the domain, recorded 2026-07-29 and reaffirmed since (PIPE_AUDIO_MUSIC §0). They are the bars every deliverable in this document serves.
2.1 The thirty-year bar
Music so nostalgic people will play it for thirty years. The exemplar lineage is broad — SNES-era, Zelda, Undertale, Sonic, Super Mario 64, Skyrim, Halo, Pokemon, Animal Crossing, Donkey Kong — scores that live outside their games. They endure on melody and leitmotif, not production sheen (PIPE_AUDIO_MUSIC §0, the thirty-year bar).
Josh expanded the bar on 2026-08-05: twelve to twenty different and unique tracks per chapter, two to seven minutes each, inspired by the cultural region and its authenticity and then made into a masterpiece, evaluated against real exemplars whose formulas are derived by measurement (MASTERPIECE_PROGRAM §0). Against seventy-nine nodes that is nine hundred and fifty to sixteen hundred tracks, and the program's answer to that arithmetic is a leitmotif economy rather than a library: a track is a motif crossed with a regional palette crossed with a purpose class, each of the three authored once and combined per chapter (MASTERPIECE_PROGRAM §1 and §5).
2.2 Melody first, and the hummability test
Hero and canonical themes are leitmotif-first. If a theme cannot be hummed after one hearing, it is not the theme yet. Themes iterate until they pass, judged with evidence like everything else (PIPE_AUDIO_MUSIC §0, melody-first and the hummability test). A theme does not proceed on schedule pressure (COMPOSER_BRIEF_PACK §8).
2.3 Per-culture, per-period scoring
Josh's verbatim direction: music is tuned to fit each culture and their time period, using instruments, chants, vocals where appropriate but mostly no vocals other than hums chants and harmonies, ambient noises. Period-appropriate instrumentation per region; vocals default wordless; ambient texture is part of the score (PIPE_AUDIO_MUSIC §0, per-culture per-period scoring).
2.4 The sourcing law
Reference-composition to a culture's musical register is permitted; sampling or reproduction of consecrated performance never is (PIPE_AUDIO_MUSIC header ruled-state line and §4). The doctrine makes the law executable rather than paralysing by splitting grammar from repertoire: grammar is free everywhere — mode and final, reciting tone, form-shape, ornament class, ensemble behaviour, drone behaviour, timbre class, tuning system, timeline — while specific transmitted repertoire, real liturgical or scriptural text, and any sampled or model-conditioned recording of consecrated performance are excluded everywhere (MUSIC_COMPOSITION_DOCTRINE §3.4).
2.5 The concert-performable orchestral bar
Full orchestras will perform this playlist at concerts, operas and Broadway. The hero-theme lane's deliverable is a composed work, not just rendered audio: real notation that a human orchestra can sit down and play. Rendered stems serve the game build; the score serves the concert hall; both come from the same composition (PIPE_AUDIO_MUSIC §0, the concert bar; MUSIC_COMPOSITION_DOCTRINE §5.1).
2.6 The theory-depth directive
Josh's verbatim direction: you will need to understand music theory in depth to create this, understand the patterns, and how everything comes together (PIPE_AUDIO_MUSIC §0, the theory-depth directive). The doctrine is the artifact that discharges it, and no theme authors without it.
§3 Music direction
3.1 Melody is the artifact
A cue is delivered when a tune survives it, not when a mix sounds good (MUSIC_COMPOSITION_DOCTRINE §0.1 first law). Three consequences bind.
One composition, two renders. The score is master; the game stems are a build of it (MUSIC_COMPOSITION_DOCTRINE §0.1 second law).
The register is the identity. Culture is carried by tuning, timeline, ensemble and performance practice, never by an exotic-signifier scale bolted onto a Western tune (MUSIC_COMPOSITION_DOCTRINE §0.1 third law).
Generation never originates identity. It carries bed and texture, inside a register card that has already declared what is lawful there (MUSIC_COMPOSITION_DOCTRINE §0.1 fourth law).
The head-cell law, binding and scoped. Every canonical theme whose motif_class is tune is built on a three-to-five-note head cell, singable inside a ninth, rhythmically distinct enough to be tapped without pitch, and either triad-outlining or purely stepwise — chosen deliberately per motif, never both. Exactly one deviant element carries identity: one unusual leap, one borrowed degree, one metric oddity. Two deviants and it stops being hummable (MUSIC_COMPOSITION_DOCTRINE §1.2). The law is scoped to motif_class tune because the roster contains identities that are deliberately not tunes, and a row with motif_class unset cannot be scored and cannot be rendered (MUSIC_COMPOSITION_DOCTRINE §1.3).
Motif class tune takes the head-cell law and the full hummability checklist.
Motif class timbre takes the atmosphere bar, never the hummability checklist. The Aquatic Ambience counter-case is a thirty-year survivor that is not a one-hearing hummable tune, and applying the wrong bar to it would reject it (MUSIC_COMPOSITION_DOCTRINE §1.3 and §2.2).
Motif class harmonic takes the progression-recognition test, which is authored and dry-run before the vril and wonder theme starts (MUSIC_COMPOSITION_DOCTRINE §1.3 and Appendix T.4).
3.2 The leitmotif architecture
Twenty-two threads do not get twenty-two independent themes; that exceeds any listener's recognition budget and produces mush (MUSIC_COMPOSITION_DOCTRINE §1.4). Three tiers.
Tier A, cardinal: exactly twelve rows, concert-scorable, the only motifs allowed a full sixteen-to-thirty-two-bar statement. The roster is pinned and Josh-approved (MUSIC_COMPOSITION_DOCTRINE §1.4 and §1.5). The live registry asserts the twelve and the applier refuses to write if the count moves in either direction (LEITMOTIF_ARCHITECTURE §1).
Tier B, family: seven families over the twenty-two threads plus the realm family, each a shared head cell and a shared harmonic signature. Inheritance is the default; independence is earned against three tests — five or more nodes, an on-screen agent, and an arc that changes state (MUSIC_COMPOSITION_DOCTRINE §1.4; LEITMOTIF_ARCHITECTURE §2).
Tier C, tags: one-to-two-bar region, boss and faction cells built from a Tier-A or Tier-B cell in the local register, never required to be hummable, only coherent with their parent (MUSIC_COMPOSITION_DOCTRINE §1.4). Any Tier-C cue whose region also carries a Tier-A payoff quotes the parent head cell at least once (§1.4, the Tier-C quotation rule).
The transformation catalogue is reserved vocabulary, not decoration. Mode shift means moral or fortune reversal; augmentation means weight, memory, ceremony, death; diminution means panic or pursuit; fragmentation means foreshadow before reveal and is the combat-layer form; reharmonisation means wonder or revelation; contour inversion means corruption or the double; metric recast means culture and era relocation; combination is the most expensive and most earned event, a new dramatic condition; closure change is the payoff itself. An operation used without a dramatic reason is drift, and a Tier-A motif budgets six to ten distinct transformations across the game (MUSIC_COMPOSITION_DOCTRINE §1.6).
Payoff architecture across dozens of hours. Withhold closure — reserve the authentic cadence on the tonic for one node in the whole game, write the resolution first, then delete it from every earlier statement. Plant diegetically — a motif the player physically inputs is retained far better than one merely heard. The single-theme arc — the Prologue statement and the Ch-77 statement are the same tune and the player must feel the distance. Recombination as reveal — motif combination states a relation with no dialogue (MUSIC_COMPOSITION_DOCTRINE §1.7).
Reveal discipline is a music gate, and it is mechanical rather than promised. Every theme row carries reveal_gate. Before its gate, only a disguised or inverted derivation of a gated motif may sound, and its kinship must be inaudible to a fresh listener (MUSIC_COMPOSITION_DOCTRINE §1.8). The live registry carries reveal gates on eleven of thirty-eight rows: four at CH_38, three at CH_76, four at CH_77. The cue emitter refuses to emit any cue whose reveal_gate is later than its own chapter if the cue name, its site or its trigger context carries a held token, the gate source fails closed on an unparseable or contradictory value, and must-fire fixtures prove both refusals fire (LEITMOTIF_ARCHITECTURE §3.4, the firewall). The antagonist pattern is the hardest case: fragmentation is the only operation permitted before Ch 38, a different local carrier every time, one statement per node, never the phrase, never a swell, never on a face (LEITMOTIF_ARCHITECTURE §3.4; COMPOSER_BRIEF_PACK §2).
3.3 The measurement instruments
The bar is held by instruments that can fail, not by assertion.
The nostalgia rubric is the exemplar-anchored scoring bar: seven axes, every band naming the exemplar track it came from, no single scalar anywhere, and what cannot be measured declared rather than dropped (NOSTALGIA_RUBRIC §1). Seven features are hard, and any hard feature failing forces its whole axis below regardless of rate (§3). It replaced proxy metrics that scored a theme against itself and therefore always reported success (§0).
The anti-plagiarism tooth is the licence position made measurable: grammar is learned, melodies are never reproduced. Candidates are reduced to transposition-invariant interval sequences and screened against a reference set held as interval sequences only, with two refusal rules and mutation proofs in both directions. The screen exists to refuse and is never a generation input (NOSTALGIA_RUBRIC §5). Its honesty clause binds: CLEAR means no seed signature matched, never clean (§5, seed-set honesty).
The exemplar corpus is the listening bar. 182 rows in build/audio/exemplars/CORPUS.json with a verified phone-first buy sheet; as of 2026-08-07 the corpus reads 153 MISSING, 19 ACQUIRED, 10 CARDED, with 1181 measured formula cards (docs/PIPELINE_LEDGER.md §7 row F1, and MUSIC CORPUS — THE STARTER TEN LANDS, 2026-08-07). The acquisition tool runs a positive control before it will proof anything and refuses rather than guesses on an unmatched file (MASTERPIECE_PROGRAM §10.1).
The corpus integrity discipline is part of the instrument, not overhead. Six rows were found pointing at the wrong audio, three were demoted with a mandatory reason, the status ladder now runs both ways, a calibrated game-agreement floor decides eligibility before title matching, and one known-false row is left standing on purpose with its defect recorded rather than guessed at (docs/PIPELINE_LEDGER.md, MUSIC EXEMPLAR CARDS and MUSIC CORPUS INTEGRITY, 2026-08-07). A seventh poison pair is open: the floor does not yet score token containment asymmetrically, so a wholesale rescan must not run until it does (MUSIC CORPUS — THE STARTER TEN LANDS, 2026-08-07).
The pattern-derivation null result binds as honestly as a positive would. The first pass across the cards found zero of one hundred and three measured axes separating the corpus rows from their own album siblings, with the honest leave-one-out composite at AUC 0.586 plus or minus 0.085 — about one standard error from chance. All four pre-registered hypotheses failed. The one usable result is an envelope usable as a rejection filter and useless as a target to optimise. The artifact says in its own words that a null result at that N is not evidence of absence (docs/PIPELINE_LEDGER.md, MUSIC PATTERN DERIVATION, 2026-08-07). The consequence for this document: no formula is asserted as the masterpiece formula, and rung three's predictor is deferred with a reason rather than a shrug.
3.4 How the melody is actually made, as measured
The 2026-08-05 melody A/B answered this by measurement rather than preference, and its finding is binding on the lane (PIPE_AUDIO_MUSIC §8).
Authored composition is the canonical-theme method's melodic source. Authored material delivered the brief's licensed leap and its head figure one hundred percent of the time, against twelve and a half percent and zero percent from the strongest text-only prompt the lane can write, over sixteen candidates. The tune is authored; it is not drawn from a prompt (§8, what is promoted).
The generator is not promoted as the arranger of that material. Repaint holds what it is handed exactly and does not continue it; cover overwrites; lego does not carry. The arrangement problem is open and named (§8, what is not promoted, and the decomposition).
A landed reading was withdrawn on re-measurement and the correction is kept on the record, not tidied away: a whole-file motif overlap that included the byte-identical held region inflated its own result, and the deciding measurement now runs on the generated region alone (§8 Q11).
The realisation rung is built and its licence is read. sfizz under BSD-2 driving VSCO 2 Community Edition and VCSL under CC0 renders the authored scores; Sonatina Symphonic Orchestra is refused on a prohibition its licence genuinely triggers; BBC SO Discover is rung two on evidence rather than schedule, with its no-conditioning clause carried onto every record (§8.1 and §8.2).
The loop seam is composed at the realisation rung, not repaired at it, and the wrap is a sum of the carried release rather than a fade (§8.2, the composed seam).
The acceptance battery did its job by failing. A first mix that read as a melody still failed the deviant-recoverability check, because octave doubling drowned the one authored leap the detector exists to find. A listening-only review would have called that mix the better one (§8.2, the battery on the mixdown).
3.5 The generated-audio posture, and the licence law
Generation owns beds, ambience and variation volume. It never originates a motif. Motif quotation in a Tier-C cue, or the motivic-residue layer of a bed, is a composed stem laid over generated material or fed as audio conditioning — never a prompt, because the generator has no melodic conditioning that could honour a specific three-to-five-note cell. The schema field motif_source makes the rule checkable rather than intended (MUSIC_COMPOSITION_DOCTRINE §6.7, the motif boundary).
The bar never routes down-tier. A bed may be generated at speed; an identity theme may not (PIPE_AUDIO_MUSIC §0, the lane split).
Inside the slice the generative lane's lawful share is small, and that is stated rather than discovered later. Across the fourteen slice nodes the register cards record generation_eligible NONE for the pitched material of CH_04 through CH_07 and CH_09 through CH_13, and PARTIAL for CH_02/03 and CH_08. The slice's music plan is budgeted on that basis (MUSIC_COMPOSITION_DOCTRINE §6.8).
Three care carve-outs are pass-or-fail on benchmark day: metre, where a non-metric or additive tradition opts out of bar-quantised vertical layering; tuning, where a non-12-TET card routes its tuned layers to the authored lane or through a declared retune pass; and ornament, where the identity is carried by gamaka, portamento, speech-glide or paired detuning (MUSIC_COMPOSITION_DOCTRINE §6.8).
The licence law is per-row, not per-document. The generating model and its licence are written into every artifact alongside the prompt hash (PIPE_AUDIO_MUSIC §4, the sourcing law operationalised). Every produced bed writes its generating model and licence into T0_SFX_Registry.generation_licence_ref, and every cue row carries licence_class alongside bed_source (LEITMOTIF_ARCHITECTURE §3.6, the licence constraint).
Two licence walls reinforce the sourcing law from the other side, and both are triggered prohibitions rather than cautious readings. The Sonniss bundle licence prohibits using the audio to train AI or ML, so no Sonniss file is ever a fine-tune corpus, a LoRA corpus, or an audio conditioning input for any generator (PIPE_AUDIO_MUSIC §4; LEITMOTIF_ARCHITECTURE §3.6). The Spitfire EULA's redistribution clause forbids derivatives usable as samples in a sampler, so no Spitfire sample may ever be a conditioning, fine-tune or LoRA input either (PIPE_AUDIO_MUSIC §8.1 option B).
Exemplar audio is measured, never fed. Analysis for understanding is ordinary use; the bytes never become model input — no training, no fine-tuning, no conditioning, no derivation. Josh ruled and settled this posture in his own words on 2026-08-05 (MASTERPIECE_PROGRAM §2 and §9).
Excluded from every shipped path on licence: MusicGen, AudioCraft and AudioGen on CC-BY-NC-4.0; Suno and Udio on indemnification and export; Stable Audio as revenue-capped fallback only (PIPE_AUDIO_MUSIC §1 demotions and exclusions; T99_Translation_Audio §2.1).
3.6 The spend gates and the publish gate
The proof-before-spend gate. Josh, verbatim: give me a few sample tracks so I can know you can actually do this and improve before I go spend $1000 more. The sample-proof lane is the gate on the remaining Class-A spend (docs/spine/DECISIONS_PENDING_JOSH.md, RULED 2026-08-06 THE BRIEF-BUNDLE ANSWERS, the music sample-proof gate). The physical disc buys wait on the music-loop proof landing (RULED 2026-08-06 THE EAR PICK IS C2, sequencing note).
The generator verdict. Josh, verbatim, after listening to the three loop-proof beds: that music is still terrible. The $900 digital bulk stays held, the music-model survey launches, and the generator itself is the candidate for replacement while the measurement rig and cards carry over unchanged whatever generates (docs/spine/DECISIONS_PENDING_JOSH.md, RULED 2026-08-06 THE MUSIC GENERATOR VERDICT).
Audio joins the publish gate. Zero exemplar tracks have passed Josh in any audio class, so the served tracks come off the site; music publishes only when its factory is complete, same as every asset class. A dedicated minimisable review surface carries only the current ungraded candidates, and a candidate turned down is removed until its successor exists (docs/spine/DECISIONS_PENDING_JOSH.md, THE REVIEW SURFACE RULES and THE FACTORY-COMPLETE PUBLISH GATE, 2026-08-06 evening).
Building never pauses. The publish gate holds publication, not work: music continues in full parallel throughout (docs/spine/DECISIONS_PENDING_JOSH.md, THE FACTORY-COMPLETE PUBLISH GATE, 2026-08-06 evening).
3.7 Adaptive delivery
The composition must survive interactivity, and two consequences bind the composition itself.
The motif carrier is per-state, not global. At every declared state exactly one stem or stem group carries the head cell, named on the state-machine row, and the mute test runs once per state rather than once per cue. That is what makes instrumental-migration escalation legal: migration changes which stem holds the role and never leaves it vacant (MUSIC_COMPOSITION_DOCTRINE §1.11 and §6.5).
Interlock-emergent registers declare an atomic stem group. In hocket, kotekan and mbira interlock the melody exists only in the composite, so the minimum carrier unit is the interlock pair or trio, that group is indivisible, and no runtime parameter may split it. Enforcing a one-stem test there would force the composer to rewrite the theme as a single-voice line — the Westernising retrofit arriving through the stem contract (MUSIC_COMPOSITION_DOCTRINE §1.11).
Silence is a first-class state, not an absence (MUSIC_COMPOSITION_DOCTRINE §6.3).
Hysteresis is mandatory and maximum de-escalation latency is a declared number per region, not an emergent one (MUSIC_COMPOSITION_DOCTRINE §6.3).
Closure-change payoffs are scripted and non-interruptible, or the engine will fade the one cadence the game has been saving (MUSIC_COMPOSITION_DOCTRINE §1.11).
Adaptive layers are built by successive additive passes on a base checkpoint, not carved out of a finished mixdown. That is what the stem benchmark measured, and the design changed to match rather than the measurement being footnoted (PIPE_AUDIO_MUSIC §6.0 Q1).
Tempo is measured from the render and never assumed from the prompt; loop-seam cleanliness is a per-artifact gate, not a property of the generator (PIPE_AUDIO_MUSIC §6.0 Q2).
§4 Sound effects
4.1 The registry is the data home
T0_SFX_Registry [ACTIVE v1.0] carries 680 rows across 41 columns and is the addressable unit of work for the lane. Addressing is already ruled and already columned: sfx_id to ue_sound_asset_path, final asset at /Game/Audio/SFX/SFX_<sfx_id>, swap by path takeover, with cue-fired telemetry and a muted-wired fixture as the standing teeth (PIPE_AUDIO_MUSIC §3).
The registry's realization state is honest and low: 662 of 680 rows carry no ue_sound_asset_path, every row reads generation_status pending, and generation_method, build_asset_path, sound_designer_in_loop_required and generation_licence_ref are empty on every row. Against that, the import chain works and 41 assets are in-engine — a six percent realization rate on a fully populated registry (docs/PIPELINE_LEDGER.md §7 row F2).
4.2 Licence posture and the authored-versus-generated split
Licensed-first sourcing. Royalty-free-owned library material is the backbone: Sonniss GDC bundles as primary, BOOM Library as the buyout fallback (PIPE_AUDIO_MUSIC §1 stack table; T99_Translation_Audio §2.1). The library is direct-use-only and can never feed a fine-tune (PIPE_AUDIO_MUSIC §4, the sourcing law operationalised).
Generative fill under an unrestricted licence. MOSS-SoundEffect v2.0, Apache 2.0, natively 48 kHz, is the primary local generative filler; ElevenLabs' SFX endpoint remains an available per-call-cost lane (PIPE_AUDIO_MUSIC §1 stack table; T99_Translation_Audio §1 reconciliation 2).
Creature vocalisation is a sound-design task, not a generative one. The 132 creature vocalisations run through the designer-in-loop lane and that carve-out survived the composer retirement explicitly (PIPE_AUDIO_MUSIC §1 stack table and §0.0).
The special-asset class takes the attended read. 43 rows carry cultural_depiction_required True and are marked cultural_authentic_audit_required; 9 rows are canonical_track; 55 rows carry hard_line_anchors. Any cue whose cultural_depiction_required or hard_line_anchors cell is non-empty, and every realm cue, takes an attended §17.1 read that cannot be delegated to a prompt. The attending party is a fresh-context critic (PIPE_AUDIO_MUSIC §0.0 and §4, unattended versus attended).
The Grand Sage is excluded from every synthesis lane. GRAND_SAGE_REVEAL_VOICE is a non-lexical vibrational-imprint sensation asset per HL_0046, ruled 2026-07-23c: it never speaks, no scene voice_ref may resolve to it, and every audio pass skips it. The gate grand_sage_silence enforces it (T0_Voice_Registry GRAND_SAGE_REVEAL_VOICE hard_line_relevance cell; PIPE_VOICE §3; harness/gates_config.json gate grand_sage_silence).
4.3 Grading posture
Zero SFX rows are graded today. The lane is exemplar-first: one SFX family — the combat impact family is the ledger's named candidate — is graded against the feel comparators before the factory runs (docs/PIPELINE_LEDGER.md §7 row F2). The impact-audio cue shape is inherited rather than invented: a two-part envelope of low-end weight resolving into a high-frequency read, with hit-stop a shared audio and animation moment, and tier escalation expressed as a low-end mass delta while the high-frequency identity read is held constant (PIPE_AUDIO_MUSIC §7.2).
Audio is wired at authoring time, not added as a post-pass. The proving map the art and animation lanes converged on carries cue audio wired at authoring time, so one surface serves screenshot and listen together (PIPE_AUDIO_MUSIC §7.1). The audio binding is declared on the entity row, where the entity is authored, rather than re-derived in a separate wiring step (§7.3).
The standing gate sfx_coverage certifies that every row is categorised or declared-waived, and reads 680 of 680 categorised with 23 of 23 fixtures at the time of writing (harness/gates_config.json; harness/sfx_coverage_scorecard.json).
§5 Foley and footstep
Foley is the domain's best-wired lane and its wired shape is the shape of done for every other audio class. In-engine today: HumanityFootstepComponent, DABP_FoleyAudioBank, ten BP_AnimNotify_FoleyEvent_* notifies covering walk, run, jump, land, scuff and hand-plant left and right, and HumanitySurfacePhysicalMaterial, with Content/Audio/Mix present (docs/PIPELINE_LEDGER.md §7 rows F3 and F2).
The data home is T0_SFX_Registry, category FOLEY.
Footstep terrain, foley costume and weapon material substrate are read from the region page's own sections rather than invented, per the sfx read-set (harness/route.py sfx).
The exemplar is one surface's full foley set, graded in-game for feel — low cost, high slice value (docs/PIPELINE_LEDGER.md §7 row F3).
Runtime variation runs on engine-native procedural DSP, unrestricted by licence (PIPE_AUDIO_MUSIC §1 stack table).
The open question this lane inherits is derivation: when mass and material are derived per object rather than authored, the impact sound has to be derived from the same two values, or every collision sounds like the same knock. This is named as the lane's top re-fetch rather than assumed solved (PIPE_AUDIO_MUSIC §7.4).
§6 Voice and VO
6.1 The canon lines
Per CVD §15.1, every option is fully written and fully voiced. The game does not use short-form voiced responses with read-only text options; all six options at every dialogue node receive the same voice-acting treatment (T3_Core_Characters [ACTIVE v1.0] §2.5). Synthetic generation operates during development; professional cast recording follows post-crowdfunding against the established biography-parameter voice profiles, with actors performing multiple takes across biography states (T3_Core_Characters §2.5).
The voice the player hears is the voice their biography parameters compose at that moment, and the register modulates continuously rather than stepwise across the five integrity bands: warm and centred at Truly Good, steady across the widest emotional range at Good Enough, neutral with the angular wisdom register at Drifting, clipped and less resonant at Evil, operational coldness at the Pure Evil terminal band — never theatrical, never villainy performance (T3_Core_Characters §2.5). Six options per beat, silence always valid where context supports it, biography options unmarked and path options subtly marked (T3_Core_Characters §2.4).
6.2 The partial-VO ship ruling, as data
The ship policy is partial VO, and it lives as a per-line column rather than as prose, because a policy that lives only in a decisions register cannot be executed or audited (T99_Translation_Audio §2.4; docs/spine/DECISIONS_PENDING_JOSH.md twenty-sixth sitting item 4).
T0_Dialogue_Line.vo_tier carries voiced_performance, synthesis_permitted or subtitle_only per line, and that column is what the VO batch reads.
The synthesis_permitted state covers named non-story classes only — barks, ambient crowd, vendor one-liners. It never covers story dialogue and it never covers a hard-line beat.
The voiced_performance state is the hero and canonical lane, behind the clone-authorization gate.
The subtitle_only state is a real shipped state, not a failure state. The subtitle floor ships regardless.
There is no runtime TTS lane in canon and this document does not create one; voice generation is bake-time only (PIPE_VOICE scope line).
Reconciling the two: CVD §15.1's fully-written-and-fully-voiced is the SHIPPED-GAME canon and remains the destination. The partial-VO ruling governs what the pre-crowdfunding build ships while professional cast recording is unfunded. The vo_tier column is the mechanism that records the distance between the two honestly per line, rather than letting a build silently claim the canon line it has not yet met.
6.3 Voice substrate posture
The flattening defect is the named failure. A single-vendor English-only roster would voice all thirteen cultural registers in the same two accents, and the registry itself forbids it (PIPE_VOICE §1). The lane's primary for culturally-registered named NPCs is a description-driven voice-design model under a permissive licence, chosen partly because generating from description rather than a reference clip makes it structurally incapable of cloning a real person (PIPE_VOICE §1 and §4).
Accents are never picked from an ethnicity menu. Three vectors only: neutral synthetic delivery with culturally registered word choice, so the register row does the cultural work and the timbre does not; description-driven voice design where every clause cites a register field; or consented community voice talent, cloned only through the authorization gate (PIPE_VOICE §4, cultural care).
The clone gate is a human seam and stays one. Cloning is never unattended, cross-cultural clone candidates route to prohibited automatically, and documented-living-tradition roles route to pure synthetic generation by default (PIPE_VOICE §4; T99_Translation_Audio §2.4).
Licence is per-row and enforced. generating_model and model_licence ride on every generated voice row, and a licence-provenance tooth fails any voice asset whose licence is not on the shippable allowlist. The lane's own worked example is why this must be per-row: a vendor's successor model flipped from permissive to non-commercial while the predecessor stayed permissive (PIPE_VOICE §3 and §1 licence kills).
Model promotion is evidence-gated and demotion is retroactive. A model enters a lane only on a same-day licence re-read, a blind A/B against the lane's reference bar, and zero findings from the accent-care lens across the register sample set; demotion re-flags every row it generated (PIPE_VOICE §4).
6.4 The natural-voice register
Every player-facing line sounds like a person, and music and voice are both player-facing surfaces (docs/NATURAL_VOICE_DOCTRINE.md §1; MUSIC_COMPOSITION_DOCTRINE §0).
Quest text is what someone said or what the protagonist thinks, never a system's summary. Requirements and tracking speak plainly. The HUD's short lines stay short and stay human (NATURAL_VOICE_DOCTRINE §1).
Names and speech carry their people. How a person speaks presents their culture — an elder's cadence, a healer's directness, a broker's oil — drawn from the persona system and T0_Voice_Registry. Boss display names become what people call the being in-world (NATURAL_VOICE_DOCTRINE §2).
Translation carries meaning, not literalism (NATURAL_VOICE_DOCTRINE §2).
The critic lens for this register is two questions: does a person say this, and does this carry the culture (NATURAL_VOICE_DOCTRINE §5).
§7 Care — CVD §17.1 at the audio layer
The care standard is CVD §17.1 and this document adds no second policy on top of it; it states the audio-layer expression of the same discipline (T99_Translation_Audio §2.7). Care means weight and authenticity. Care never means omission, reverence-gloss or rated-G sanitization (CVD §17.1). Care-tier inflation is itself a defect and the timid option is usually the wrong one (CVD §17.1 calibration paragraph). The hard-line floor is the only absolute and never loosens (CVD §17.1 calibration paragraph).
The collective protection, expressed in music, has exactly one failure mode. No region's own register is used as the threat signal, systematically darkened, inverted or menace-coded. The darkening operations belong to the antagonist material and travel with it, never with the place (LEITMOTIF_ARCHITECTURE §4; MUSIC_COMPOSITION_DOCTRINE §1.9).
The corollary is equally binding and points the other way. An individual antagonist may be scored fully inside his own culture's register, with the same craft as that region's protagonist cues. A blanket register ban is the named failure, not the safe option — it would produce the stranger result that every villain sounds Western-orchestral while the local register is reserved for scenery (LEITMOTIF_ARCHITECTURE §4, the corollary; MUSIC_COMPOSITION_DOCTRINE §1.9). Care-tier evasion is a defect and so is care-tier inflation.
Per-motif allow and deny register lists are required fields, and the deny list carries the region-as-threat check. A motif with an empty deny list has not been reviewed (MUSIC_COMPOSITION_DOCTRINE §1.9). The live T0_Theme_Registry carries allow_register_list and deny_register_list as real columns.
One rule for every sacred tradition, not one rule per continent. Treating Gregorian chant as free public grammar to invent within while barring an equally documented liturgical tradition even from original composition is the colonial register arriving through timidity, and the doctrine corrected it: every tradition in the slice gets a lawful composed vocal lane, or its card states in writing which named exception removed it and why (MUSIC_COMPOSITION_DOCTRINE §3.4).
Named sacred exclusions are named, per tradition, in the tradition's own words. The doctrine's per-culture register cards carry them by name rather than by a general sentence about sacred music, and the brief pack copies each page's own prohibition rather than paraphrasing it (MUSIC_COMPOSITION_DOCTRINE §4.2 through §4.12; COMPOSER_BRIEF_PACK §7).
Converting a living people into ambience is a recognised colonial move and is refused. Scoring a people out of the frame while keeping their forest is a care failure; endangerment is an argument for accuracy and consultation, not for substituting birdsong (MUSIC_COMPOSITION_DOCTRINE §4.5, the Vedda correction).
No living people's idiom scores a deep-time stratum. Scoring the Paleolithic with a present-day Indigenous register is the primitive-coding this discipline exists to prevent, and it is also false. Deep-time strata are scored with harmonic-series material, breath and struck-stone timbres, and the adjacent living register is explicitly barred from them (MUSIC_COMPOSITION_DOCTRINE §4.0.3).
Period defects are care defects, not taste notes. Scoring Flores with gamelan colour, or Angkor-era Bali with a twentieth-century ensemble, or a 1200-to-1400 Swahili chapter with late-nineteenth-century taarab, is the same class of error as scoring Wales with a sitar (MUSIC_COMPOSITION_DOCTRINE §4.2, §4.3 and §4.7).
The generic wash is a care failure, not a taste note. The word world-music does not appear as a descriptor, per the CVD §17.1 banned-descriptor line, and a generic sacred wash under a cardinal theme is a failure of this standard (MUSIC_COMPOSITION_DOCTRINE §1.9 and §3.5).
The register must be recognisable to someone from that culture, and that is a claim only a person from that culture can settle. Ship-tier register review by a named tradition-bearer or specialist gates public release of that register's music, with the reviewer's objections recorded verbatim including any overruled and why. The solo-executable interim tier is a specialist-literature verification pass, labelled proxy-tier and never claimed as the real thing (MUSIC_COMPOSITION_DOCTRINE §4.0.1).
The audio expression of collective protection carries into voice. The care review listens for caricature markers, flattening measured as speaker-embedding distance, villain-accent correlation, and care-tier inflation — where a voiced hero set skewing Western while culturally registered chapters stay subtitle-only would carry a bias the strings do not. Hero-lane VO coverage is checked against the register map, not just the node list (PIPE_VOICE §4, the care review).
Protected property is referenced, never reproduced or staged as spectacle (CVD §17.1, protected property referenced not reproduced), and no cue name, track title, subtitle, achievement or soundtrack listing may carry a held reveal token before its gate (COMPOSER_BRIEF_PACK §2).
§8 Production, evidence and QA
8.1 The exemplar-first law
An audio class runs its factory only after one exemplar of that class has passed. Zero graded exemplars exist in any audio class today: F1 music records 20 staged tracks, none graded; F2 SFX records 0 graded; F3 foley records 0 graded (docs/PIPELINE_LEDGER.md §7 rows F1-F3). The named exemplars are one hero theme to final-candidate against the exemplar corpus, one SFX family graded against the feel comparators, one surface's full foley set graded in-game for feel, and one six-option dialogue beat fully voiced and played in-engine (docs/PIPELINE_LEDGER.md §4 row C9 and §7 rows F1-F3).
8.2 The honest tier register
Every claim about audio quality carries its tier, and the tier does real work rather than decorating a report.
Structure, greybox-feel and final are the three honest tiers for play-quality claims generally.
The music lane's own tiers, in the order they were earned: GENERATED SCORE — ITERATION for anything the generation ladder emits (PIPE_AUDIO_MUSIC §0.0); AUTHORED_HEAD_CELL for an authored score in a sketch realisation (§8, the honest tier); AUTHORED_SCORE_ORCHESTRAL_RUNG_ONE for the sampled realisation, served with a note stating outright that it is a public-domain community sample library rather than a professional scoring session (§8.2, what is promoted).
Notation emitted by extraction from a render rather than authored is tiered EXTRACTED-NOT-AUTHORED, and the orchestral pass that would make it concert-ready is a named later rung rather than a claim (COMPOSER_BRIEF_PACK §8).
No concert-performable claim is made until the concert definition of done closes for that theme. Orchestrable is not orchestrated and read (MUSIC_COMPOSITION_DOCTRINE §5.9).
A tier that is not served is not landed. A build that serves no audio may not describe itself as clean (PIPE_AUDIO_MUSIC §8.2, the site finding).
8.3 Critic gates and instrument honesty
Every quality claim carries evidence: captures, score sheets, or a named reviewer with a date. Nothing is signed off on description (MUSIC_COMPOSITION_DOCTRINE §7.8).
Listen-checks run on captured playback by a fresh-context critic, never on a described claim — the audio analogue of the visual-verification directive (PIPE_AUDIO_MUSIC §4, the quality gate).
The reveal check runs fresh-context and capture-backed on every pre-gate statement (MUSIC_COMPOSITION_DOCTRINE §1.8, gate 8).
Gates that cannot be run are gates everybody ignores. Every human dependency carries an owner, a trigger milestone, a cost estimate, and a solo-executable interim proxy labelled proxy-tier and never claimed as the real thing (MUSIC_COMPOSITION_DOCTRINE §0.1 sixth law, §5.9 and Appendix T.1).
An instrument that cannot fail is not an instrument. Extractors self-test against material with a known answer; a rubric that cannot fail a bad melody is a formality, and the hook-less control is what proves it is not (MASTERPIECE_PROGRAM §2, instrument controls; NOSTALGIA_RUBRIC §8).
The standing gate suite carries twelve audio teeth, each exit-code-bearing on every run: grand_sage_silence, sfx_coverage, nostalgia_rubric, music_acquire, music_battery, music_cards, music_compose_arm_b, music_orchestrate, music_picks, music_recompose_realise, music_seams and music_stack (harness/gates_config.json).
Every artifact carries a paired licence record and generation record, and the pairing is enforced by a check with a positive control rather than asserted (PIPE_AUDIO_MUSIC §8.2, the pairing law).
8.4 Definition of done, split by what it gates
The split exists because a set of unreachable concert-hall requirements would otherwise block the GAME build of every hero theme in a solo pre-revenue project (MUSIC_COMPOSITION_DOCTRINE §5.9).
Gates the game build, and is entirely solo-executable: notation authored on the real orchestra template with written ranges honoured; a proofreading pass clean; a self-check on range, breath, divisi and page density; the realisation rendered with sectional stems delivered at the format contract; a critic listen-check on captured playback including transitions; the registry row complete with register-card reference, reveal_gate and motif_class; and the composition checklist passed at the declared tier with the capture archived.
Gates the concert release only, and never blocks a game build: professional proofread of score and parts; the print specification; a playability read by a real orchestral player; the phonetic screen for any invented syllable, which also gates the game build if a vocal capture ships; the choir session at declared forces; and the tradition-bearer register review, which gates public release of that register's music in either lane.
8.5 Format and integration contract
Generate at 48 kHz, deliver 16-bit; the engine converts every accepted source format internally, so source bit depth is never a shipped-quality lever (PIPE_AUDIO_MUSIC §1 UE reality check and §3).
Voice masters at 48 kHz 24-bit mono; engine import at 48 kHz 16-bit mono; cloud returns are resampled at ingest, never at import (PIPE_VOICE §3).
Adaptive cues deliver one loopable file per stem role plus tempo and bar length, with clean loop points and no baked fades, because a baked fade destroys the quantization boundary (PIPE_AUDIO_MUSIC §3, the stem and loop contract).
Music addressing: T0_Theme_Registry.theme_id to ue_sound_asset_path, assets at /Game/Audio/Cues/<cat>/A_<theme_id>. SFX addressing: T0_SFX_Registry.sfx_id to ue_sound_asset_path, assets at /Game/Audio/SFX/SFX_<sfx_id>. Voice addressing: T0_Voice_Registry.voice_id to ue_sound_asset_path, assets at /Game/Audio/Voice/<cat>/V_<voice_id> (PIPE_AUDIO_MUSIC §3; PIPE_VOICE §3).
Every generated cue, line and bank binds to an id. No orchestration agent composes a prompt from a freehand brief; every descriptor traces to a named field on one of the audio registries or the region page substrate that feeds them (T99_Translation_Audio §2.5).
§9 The state of the domain, recorded honestly
Stated so that no later reader mistakes a populated registry for a realized one.
Music: 38 theme rows, 1 motif row, 20 staged tracks, zero graded. The music cue table has no in-engine consumer and there is no music DataTable (docs/PIPELINE_LEDGER.md §7 row F1).
Foley: wired and complete-shaped, zero graded (docs/PIPELINE_LEDGER.md §7 row F3).
Voice: 23 rows, no generated VO line exists anywhere under the build tree, no VO playback path in engine (docs/PIPELINE_LEDGER.md §4 row C9).
The dispatch-key blocker PIPE_VOICE recorded against T0_Voice_Registry — a duplicated header column and unpopulated row_class values — is REPAIRED as of this writing: the file carries 21 distinct columns, every row is 21 fields, and row_class reads 13 cultural_register plus 10 tts_casting (PIPE_VOICE §4 blocker, against the live file). The dossier's blocker paragraph is now stale and is named in §10.
The head-of-chain blocker PIPE_AUDIO_MUSIC recorded for the SFX and music lanes — the region page's machine cue table not existing, leaving the lane with no addressable unit of work — requires re-verification against the current Flores page before it is repeated as live (PIPE_AUDIO_MUSIC §4, the head-of-chain blocker).
§10 Open items flagged to the director
Nothing in this section is absorbed into the body above. Each is either genuinely new, or a live contradiction between sources that this document is not authorised to rule.
10.1 Composer-in-loop: two rulings on record, one day apart, pointing opposite ways
This is the sharpest contradiction in the domain and it is a Josh-level fork, not a drafting choice.
On 2026-08-05 Josh retired composer-in-loop in his own words: "Humans aren't composing the music… what the fuck. The entire music pipeline is yours." The dossier records the retirement as supreme and annotates exactly one row (PIPE_AUDIO_MUSIC §0.0). COMPOSER_BRIEF_PACK re-targeted itself the same day (§0), and LEITMOTIF_ARCHITECTURE recorded the consequence: the melodies now exist in notes, composed by the factory (§5).
On 2026-08-06, two separate entries in the decisions register close with "leitmotifs stay composer-in-loop" (docs/spine/DECISIONS_PENDING_JOSH.md, RULED 2026-08-06 THE BRIEF-BUNDLE ANSWERS) and "leitmotifs composer-in-loop regardless of generator" (RULED 2026-08-06 THE MUSIC GENERATOR VERDICT). Both sentences are the recording seat's summary line; neither of Josh's two verbatim quotations in those entries mentions who composes.
What this document carries, and why. Both readings agree on the load-bearing half: an identity theme is COMPOSED, never prompted out of a generator, and the bar never routes down-tier. The 2026-08-05 measurement landed on exactly that — authoring is the melodic source, generation is beds and ambience (PIPE_AUDIO_MUSIC §8, what is promoted and what is not). §3.4 above therefore states the composed-not-prompted rule, which both rulings satisfy, and does not state who holds the pen. That is the open half.
The fork for the director: is composer-in-loop retired (the 2026-08-05 verbatim), or does it stand and the factory's authored composition is an interim (the 2026-08-06 summary lines)? The practical difference is real — a human composer is a named cost, a booking trigger and an owner in the definition of done, and the 2026-08-06 generator verdict plus the held $900 bulk spend make the question live rather than academic.
RULED — CLOSED (Josh, 2026-08-08, verbatim): "Remove the human composer. We are designing that. Don't hate on that. It's another pass or few layers and patterns we may just need to figure out." The 2026-08-05 retirement STANDS and is permanent; the 2026-08-06 "composer-in-loop" summary lines were the recording seat's phrasing, never Josh's, and are superseded. From this ruling forward, "composer_in_loop" wherever it appears as a cue source class means THE FACTORY'S OWN COMPOSITION LADDER — the pass system (pass 3's development grammar, pass 4's repairs, and the further layers and patterns the program designs) — never a human hire. No booking trigger, no human-composer cost line, anywhere. The bars are unchanged: themes are COMPOSED never prompted, the score is the master, Josh's ear grades every round, and the register never frames the in-house ladder as an interim awaiting a human — designing the composer IS the program.
10.2 The concert deliverable and the retirement do not fully reconcile
The concert-performable bar asks for a composed work an orchestra can read, and the doctrine's two-deliverable law makes the score the master and the render a build of it (MUSIC_COMPOSITION_DOCTRINE §5.1). Under the generation lane the order inverts: the render is authored first and the notation is extracted from it (COMPOSER_BRIEF_PACK §8). The authored method has since moved the lane back toward the concert bar — an authored score is already the thing an orchestra could read (PIPE_AUDIO_MUSIC §8, the next lever) — but no orchestral MusicXML pass exists and the concert definition of done has never been opened for any theme. Flagged, not resolved.
10.3 Doctrine gate 5 is met only in part, and the outstanding half is owed
Plant node and diegetic surface now exist as columns on the live T0_Theme_Registry but are deliberately unpopulated, on the grounds that a diegetic surface exists in a node and must be assigned by a lane that has read that node. A first render has already landed and been served with half its gate unrepresented (LEITMOTIF_ARCHITECTURE §5, doctrine gate 5). Owed before the next render of any Tier-A row: a plant node and a diegetic surface per Tier-A theme, assigned from the spine entry it plants in, with the protagonist theme first.
10.4 The recurrence anchors are unpopulated by design and the firewall is unarmed until they land
The antagonist pattern's pre-gate statements are deliberately not placed, because a hand-typed anchor list would be exactly the invented-canon failure the repo gates against. The reveal firewall arms itself the moment the first anchor lands (LEITMOTIF_ARCHITECTURE §3.4, the recurrence anchors). Until then the firewall's teeth exist and have nothing to bite. Named so nobody reads the armed emitter as an armed gate on real content.
10.5 T0_Motif_Registry at one row
One motif row under a melody-first thirty-year bar is the domain's sharpest data gap (docs/PIPELINE_LEDGER.md §7 row F1). Whether the motif bank of forty to sixty motifs the masterpiece program calls the highest-value authoring in the whole program (MASTERPIECE_PROGRAM §7, rung 4) lands in T0_Motif_Registry or is absorbed into T0_Theme_Registry's Tier-B and Tier-C rows is an unanswered registry-shape question. It is a canon-adjacent act and is named here rather than decided.
10.6 Two dossier statements are stale against the live tree
PIPE_VOICE's dispatch blocker (PIPE_VOICE §4) describes a duplicated header column and ten classless rows in T0_Voice_Registry. The live file does not carry that defect. The dossier paragraph needs a supersession note so a later lane does not chase a repaired bug.
PIPE_AUDIO_MUSIC's head-of-chain blocker (PIPE_AUDIO_MUSIC §4) states that the region page's machine cue table does not exist and that the audio lane therefore has no addressable unit of work. A machine cue table now exists at build/audio/music_cue_table.csv per COMPOSER_BRIEF_PACK §8, and the ledger records a 216-row music cue table with no in-engine consumer (docs/PIPELINE_LEDGER.md §7 row F1). Whether the blocker is closed, partially closed, or moved is not something this document can settle from the sources.
10.7 The registry-extension conformance gap on generated audio
T0_SFX_Registry carries generation_status, generation_prompt_hash, last_generated_timestamp and regeneration_trigger as columns, but generation_method, sound_designer_in_loop_required, build_asset_path and generation_licence_ref are empty on all 680 rows, and generation_status reads pending on all 680. The five-state generation lifecycle and the licence law therefore have columns but no values. This is population, not schema, and it belongs in a batched registry pass with one fidelity re-baseline in the same commit (PIPE_AUDIO_MUSIC §3, the two verified state facts).
10.8 The music exemplar corpus carries one known-false row and one open matcher defect
EX_014 stands ACQUIRED against audio from the wrong franchise, deliberately, with its own note saying the question is answered by buying the album and looking rather than by the tool (docs/PIPELINE_LEDGER.md, MUSIC CORPUS INTEGRITY, 2026-08-07). The seventh poison pair — a row whose game tokens are a strict subset of a file's — is open, and a wholesale rescan must not run until the game floor scores containment asymmetrically (docs/PIPELINE_LEDGER.md, MUSIC CORPUS — THE STARTER TEN LANDS, 2026-08-07). Both are recorded on the corpus's own defect register; both are carried here because they bound what the measurement rig may claim.
10.9 The Care-Doctrine benchmark question is still unrun
The benchmark's cultural-substrate question — do region prompts produce plausibly registered instrumentation, or a generic wash — was never run, because all twelve region cells are correctly held pending the elevated-care read. Nothing in the landed benchmark evidence may be cited on it (PIPE_AUDIO_MUSIC §6 item 6 and §6.0's scope line). This is the domain's largest open care gate.
10.10 No dedicated audio comparator research has ever been run
Everything the comparator program contributed to this domain was routed from lanes looking at something else. Four questions a real pass would answer remain unanswered: how shipped AAA titles structure an adaptive music system's authoring surface, how SFX libraries are organised and versioned at scale, how VO is routed and QA'd against script churn, and how mix and loudness are validated automatically. The thirty-year melodic bar in particular has no comparator anchor (PIPE_AUDIO_MUSIC §7.6). Boarded, not done.
10.11 Beyond-floor items this document proposes, named for veto
Per the standing law that Josh's comments are the floor and the loop expands beyond it, three items below are this document's own additions rather than restatements of a ruling.
This document exists at all. Lifting T1_Audio_Spec out of stub state is VETOABLE 3 in the pipeline ledger — a structural finding the ledger surfaced, not an item Josh named (docs/PIPELINE_LEDGER.md §11, VETOABLE 3).
The four audio surfaces are held as one domain under one T1. The alternative is four owners. The region-page template's own consumption pairing is the evidence for the consolidation (T99_Translation_Audio §1), but the choice is this document's.
§6.2's reconciliation of the fully-voiced canon line with the partial-VO ruling — reading the first as the shipped-game destination and the second as the pre-crowdfunding build's honest state, with vo_tier as the per-line record of the distance — is this document's reading. No ruling states it in those terms, and both sources are quoted rather than paraphrased so the reading can be checked against them.
Generated by harness/site/structure_site.py — the URL path is the repo path. review root